Sample Bounded Distributed Reinforcement Learning for Decentralized POMDPs

نویسندگان

چکیده

Decentralized partially observable Markov decision processes (Dec-POMDPs) offer a powerful modeling technique for realistic multi-agent coordination problems under uncertainty. Prevalent solution techniques are centralized and assume prior knowledge of the model. We propose distributed reinforcement learning approach, where agents take turns to learn best responses each other’s policies. This promotes decentralization policy computation problem, relaxes reliance on full problem parameters. derive relation between sample complexity response error tolerance. Our key contribution is show that could grow exponentially with horizon. empirically even if requirement set lower than what theory demands, our approach can produce (near) optimal policies in some benchmark Dec-POMDP problems.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Sample Bounded Distributed Reinforcement Learning for Decentralized POMDPs

متن کامل

Pruning for Monte Carlo Distributed Reinforcement Learning in Decentralized POMDPs

متن کامل

Solving Finite Horizon Decentralized POMDPs by Distributed Reinforcement Learning

متن کامل

Bounded Dynamic Programming for Decentralized POMDPs

Solving decentralized POMDPs (DEC-POMDPs) optimally is a very hard problem. As a result, several approximate algorithms have been developed, but these do not have satisfactory error bounds. In this paper, we first discuss optimal dynamic programming and some approximate finite horizon DEC-POMDP algorithms. We then present a bounded dynamic programming algorithm. Given a problem and an error bou...

متن کامل

Optimizing Memory-Bounded Controllers for Decentralized POMDPs

We present a memory-bounded optimization approach for solving infinite-horizon decentralized POMDPs. Policies for each agent are represented by stochastic finite state controllers. We formulate the problem of optimizing these policies as a nonlinear program, leveraging powerful existing nonlinear optimization techniques for solving the problem. While existing solvers only guarantee locally opti...

متن کامل

ذخیره در منابع من

ذخیره در منابع من قبلا به منابع من ذحیره شده

{@ msg_add @}

با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

ژورنال

عنوان ژورنال: Proceedings of the ... AAAI Conference on Artificial Intelligence

سال: 2021

ISSN: ['2159-5399', '2374-3468']

DOI: https://doi.org/10.1609/aaai.v26i1.8260